Loosely inspired by the brain
The term "neural network" comes from neuroscience. In 1943, Warren McCulloch and Walter Pitts published a mathematical model of how biological neurons might compute. In 1958, Frank Rosenblatt built on this to create the perceptron, the first trainable artificial neuron. These are the historical roots of modern deep learning.
That said, an important clarification for anyone who wants to be accurate: artificial neural networks are only loosely inspired by biology. A real biological neuron is enormously complex, with thousands of chemical signals, temporal dynamics, and structural properties that no artificial model comes close to replicating. When researchers say "inspired by the brain," they mean the basic idea of connected units that signal each other, not a faithful simulation of neuroscience.
So set aside the biology. An artificial neuron is a mathematical function. Understand it as maths and everything will make sense.
"Deep learning is a class of machine learning algorithms that use multiple layers to progressively extract higher-level features from raw input."
LeCun, Bengio and Hinton, Nature 2015 — the paper that defined the field for a generationWhat one artificial neuron actually does
A single artificial neuron takes a set of inputs, multiplies each one by a weight, adds them all together along with a constant called the bias, and then passes the result through an activation function. That is the complete description of a neuron's computation.
The neuron computes z (the weighted sum plus bias), then applies the activation function f to get the output. Without the activation function, you could collapse any number of layers into a single multiplication and the network would have no more power than a basic linear model.
The weights determine how much each input contributes. A large positive weight means "when this input is high, the output should be high." A negative weight means "when this input is high, the output should be low." The bias is an offset that lets the neuron fire even when all inputs are zero, making the model more flexible.
These weights and the bias are the parameters of the neuron. Before training they are random. After training, they encode the pattern the network learned.
Activation functions: introducing non-linearity
The activation function is what makes neural networks powerful. Without it, stacking multiple layers of neurons would be mathematically identical to having just one layer. No matter how deep your network, it could only learn linear relationships, which would make it no better than linear regression.
By applying a non-linear function after each summation, you allow the network to learn curved, complex, non-linear relationships between inputs and outputs. This is the key insight.
Before ReLU, networks used sigmoid and tanh activations throughout. These functions both "squash" large values, which means gradients (the signals used for learning) become extremely small as they flow back through many layers. The network stops learning effectively. ReLU avoids this because for positive inputs, the gradient is always exactly 1, so signals can flow freely through many layers. The widespread shift to ReLU around 2012 was a major reason why very deep networks became trainable for the first time.
Layers: organising neurons into a network
One neuron is not very useful on its own. The power comes from connecting many neurons together in layers, with the output of one layer feeding into the input of the next. This creates a feedforward neural network, also called a multilayer perceptron (MLP).
Every neuron in each layer is connected to every neuron in the next layer (this is called a "fully connected" or "dense" layer). The input layer simply receives the data. The output layer produces the final prediction. The hidden layers do the work of learning intermediate representations.
The number of layers and the number of neurons per layer are both hyperparameters that you choose before training. A network with two or more hidden layers is commonly called a deep neural network, which is where the term "deep learning" comes from. Depth is not a precise technical threshold; it is a term that reflects the shift in the field toward architectures with many layers.
The forward pass: how data flows through the network
When you feed a data point into a trained neural network to get a prediction, the computation that happens is called the forward pass. It is called "forward" because information flows in one direction: from the input layer, through each hidden layer in sequence, to the output layer. Nothing goes backwards during inference.
Think of a large company processing a job application. The document first reaches the HR team who extract key facts (education, years of experience). Their summary goes to the hiring manager who assesses fit for the role. That assessment goes to the department head who makes a final recommendation. Each layer processes what the previous layer passed on, adding a higher level of interpretation. The neural network does the same thing with numbers.
What makes neural networks so powerful
In 1989, mathematician George Cybenko proved something remarkable: a neural network with just one hidden layer containing enough neurons can approximate any continuous mathematical function to any desired degree of accuracy. This result, known as the Universal Approximation Theorem, tells us that the architecture is not the limiting factor. A sufficiently large network can, in principle, learn any pattern that exists in data.
This does not mean "any network learns anything." The theorem tells us about theoretical capacity, not about whether training will actually find the right weights, or whether you have enough data, or whether the network will generalise. But it does explain why the architecture is so widely applicable: image recognition, language translation, game playing, weather forecasting, protein structure prediction. One architecture, tuned differently, does all of it.
Your first neural network in Keras
Keras is the standard high-level interface for building neural networks. It ships as part of TensorFlow and is the fastest way to go from idea to working model. The API maps directly onto the concepts you just learned: you stack layers, specify their size and activation, then compile and train.
import numpy as np from tensorflow import keras from tensorflow.keras import layers from sklearn.datasets import load_breast_cancer from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler # Load and prepare data X, y = load_breast_cancer(return_X_y=True) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y ) # Scale inputs: neural networks are sensitive to feature scale scaler = StandardScaler() X_train = scaler.fit_transform(X_train) X_test = scaler.transform(X_test) # Build the network model = keras.Sequential([ layers.Dense(64, activation='relu', input_shape=(X_train.shape[1],)), layers.Dense(32, activation='relu'), layers.Dense(1, activation='sigmoid') ]) # Compile: choose loss function and optimiser model.compile( optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'] ) # Train history = model.fit( X_train, y_train, epochs=50, batch_size=32, validation_split=0.15, verbose=0 ) # Evaluate on the held-out test set loss, accuracy = model.evaluate(X_test, y_test, verbose=0) print(f"Test accuracy: {accuracy:.4f}")
Notice a few things in the code. The network has an input layer (defined by input_shape), two hidden layers with ReLU activation, and one output neuron with sigmoid (because this is binary classification: benign or malignant). The loss function is binary cross-entropy, the standard choice for binary classification. The optimiser is Adam, a modern gradient descent variant that adapts the learning rate automatically and works well as a default.
Also notice the StandardScaler applied to the inputs. Neural networks are sensitive to the scale of their inputs in a way that decision trees and random forests are not. Features with large numeric ranges can dominate the gradient updates and make training unstable. Scaling all features to have mean zero and standard deviation one is standard practice before feeding data into a neural network.
This is not optional. If your features have very different scales (for example, age in the range 18-80, and income in the range 20,000-500,000), the income feature will produce gradients hundreds of times larger than the age feature. The network will pay almost no attention to age during training. StandardScaler (or MinMaxScaler for inputs bounded between 0 and 1) fixes this.